Papers with natural language datasets

4 papers
Reducing Gender Bias in Word-Level Language Models with a Gender-Equalizing Loss Function (P19-2)

Copied to clipboard

Challenge: Existing methods to reduce gender bias in natural language datasets are inadequate.
Approach: They propose a loss function modification approach which equalizes the probabilities of male and female words in the output.
Outcome: The proposed approach outperforms existing methods in several aspects, especially in reducing gender bias in occupation words.
Hidden Schema Networks (2023.acl-long)

Copied to clipboard

Challenge: Existing models that encode rich semantic and syntactic content are biased, but they are effective at encoding symbolic representations.
Approach: They propose a neural language model that enforces explicit relational structures which allow for compositionality onto the output representations of pretrained language models.
Outcome: The proposed model can encode sentences into sequences of symbols and infer the posterior distribution of the model from natural language datasets.
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance? (2024.emnlp-main)

Copied to clipboard

Challenge: Recent large language models (LLMs) have demonstrated remarkable generalization abilities in mathematics and reasoning tasks.
Approach: They pre-trained decoder-based language models from scratch using ten programming languages and three natural language datasets.
Outcome: The proposed models outperform natural languages on logical reasoning tasks.
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective (2026.acl-long)

Copied to clipboard

Challenge: Prior studies comparing FT and ICL have yielded mixed and inconclusive results due to inconsistent experimental setups.
Approach: They propose a formal language learning task with precise language boundaries, controlled string sampling, and no data contamination to enable a rigorous comparison.
Outcome: The proposed task offers precise language boundaries, controlled string sampling, and no data contamination.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations